Skip to main content
This guide provides detailed performance benchmarks and optimization techniques for PVAC-HFHE based on real measurements.

Performance overview

PVAC-HFHE excels at shallow circuits and scalar operations. From benchmark data:

Core operations

All benchmarks from benchmarks/README.md running on DigitalOcean Premium AMD 8-core 2.0GHz with g++ -O3 -march=native.

Key generation performance

From benchmarks/README.md:191-200:

Optimization tips

Current implementation is an unoptimized proof-of-concept. Key generation is 22x slower than BFV but only runs once per session.
Mitigation strategies:
  1. Cache keys: Generate once, serialize to disk
  2. Precompute powers: The powg_B table is the bottleneck
  3. Parallel generation: H matrix generation can be parallelized

Encryption performance

Single value encryption

From benchmark data:
  • Time: 84.11ms (mean)
  • Stddev: 2.08ms
  • vs BFV: 8x slower
  • vs CKKS: 3.6x faster
Optimization:

Depth hint optimization

From include/pvac/ops/encrypt.hpp:732-738:
Best practice:
  • Use depth 0 for additions only
  • Use depth 1-2 for shallow multiplications
  • Use depth 3+ only when necessary
Profile your circuit depth first, then use the minimum required depth hint to minimize encryption overhead.

Addition performance

From benchmark data:
  • Time: 0.012ms (12 microseconds)
  • vs BFV: 10x faster
  • vs CKKS: 87x faster

Why so fast?

From include/pvac/ops/arithmetic.hpp:165-188, addition is pure graph concatenation:
No field operations, no PRFs, just memory operations. Exploit this:

Multiplication performance

From benchmark data:
  • Time: 2.47ms (mean)
  • vs BFV shallow: 2.9x faster (7.23ms)
  • vs BFV leveled: 7.4x faster (18.28ms)
  • vs CKKS: 14.3x faster (35.23ms)

Depth performance

From benchmarks/README.md:88-98:
Exponential degradation beyond depth 2.

Optimization strategies

1. Minimize depth

2. Use ct_square for x²

From include/pvac/ops/arithmetic.hpp:227-255:
Savings: ~40% fewer product layers.

3. Tune S parameter

The S parameter controls edges per product layer. Larger S increases time/size but improves noise distribution.

Dot product performance

From benchmarks/README.md:116-124:
Implementation:
Complexity:
  • n encryptions of a: n × 84ms
  • n encryptions of b: n × 84ms
  • n multiplications: n × 2.47ms
  • n additions: n × 0.012ms (negligible)
  • Total: ~168n ms for PVAC vs ~2300n ms for BFV

Polynomial evaluation

For f(x) = 3x³ + 2x² + 5x + 7: From benchmarks/README.md:128-136:
Optimized implementation:
Saves multiplications vs naive expansion.

Ciphertext size optimization

From benchmarks/README.md:76-86:

Compaction

Automatic edge compaction when budget exceeded:
Manual compaction:
Compaction is expensive (O(E × B)) but can reduce ciphertext size by 50-80% by merging edges.

Parallel throughput

From benchmarks/README.md:175-181:
Parallel multiplication:
Speedup: ~7.4x on 8 cores.
PVAC-HFHE parallelization is coarse-grained (operation-level). RLWE SIMD is 146x faster for fine-grained vectorization.

Comparison: PVAC vs bit-level FHE

From benchmarks/README.md:32-39:
64-bit multiplication estimate: From benchmarks/README.md:150-160:
This comparison is for demonstration only. Bit-level FHE solves different problems (arbitrary boolean circuits) vs PVAC (arithmetic circuits).

Memory usage

Estimated memory for different operations:
For memory-constrained environments, use depth 0-2 operations and compact ciphertexts frequently.

Benchmarking your code

From examples/basic_usage.cpp:246-265:

Compiler optimization flags

From benchmarks/README.md:274:
Critical flags:
  • -O3: Maximum optimization
  • -march=native: CPU-specific instructions (SIMD, AES-NI)
  • -fopenmp: Parallel support
Without -march=native, performance may degrade by 30-50% due to missing PCLMUL instructions for field arithmetic.

Next steps

Depth management

Master circuit depth optimization

Arithmetic operations

Learn efficient operation patterns